Papers by Vineeth N. Balasubramanian
Mind’s Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluations of multimodal large language models (MLLMs) have demonstrated compelling visual understanding in recent years. |
| Approach: | They propose a multimodal large language model with eight visuo-cognitive tasks inspired by classic human intelligence tests organized under a novel A–R–T taxonomy: Abstraction, Relation, and Transformation. |
| Outcome: | The proposed frameworks are based on eight visuo-cognitive tasks inspired by human intelligence tests and organized under a novel A–R–T taxonomy: Abstraction, Relation, and Transformation. |
Mitigate One, Skew Another? Tackling Intersectional Biases in Text-to-Image Models (2025.findings-emnlp)
Copied to clipboard
Pushkar Shukla, Aditya Chinchure, Emily Diana, Alexander Tolbert, Kartik Hosanagar, Vineeth N. Balasubramanian, Leonid Sigal, Matthew A. Turk
| Challenge: | a new tool for analyzing and quantifying bias interactions in text-to-image models is being developed . a bias in text models can be deeply interrelated, but measuring such effects quantitatively remains a challenge. |
| Approach: | They propose a tool to quantify bias interactions in text-to-image models by analyzing and quantifying bias interactions along bias axes. |
| Outcome: | a new tool analyzes and quantifies bias interactions in text-to-image models . estimates show strong correlations with observed post-mitigation outcomes . |
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities (2025.acl-long)
Copied to clipboard
| Challenge: | Vision-language Models have been shown to be highly capable but lacking basic visual understanding skills. |
| Approach: | They propose to examine the limitations of vision-language models on visual tasks by constructing a series of tests that probe which components of design may be lacking. |
| Outcome: | The proposed tests compare VLMs to other models on visual encoders, intermediate vision-language projection and LLM-decoder outputs. |
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs (2026.acl-short)
Copied to clipboard
| Challenge: | Existing multimodal reasoning models lack generalized spatial intelligence, a new study shows . a critical gap exists in the field of vision-centric reasoning, the authors argue . |
| Approach: | They evaluate 16 multimodal reasoning models using Chain-of-Though (CoT) based thinking . they find that CoT prompting consistently degrades performance in visual spatial reasoning . |
| Outcome: | The proposed model hallucinates visual details from textual priors even when the image is absent. |